Skip to main content

GRU: The 2-in-1 Shampoo

LSTMs are powerful, but they have a problem: they are mathematically expensive.

Because they maintain two completely separate roads (the VIP Highway and the Local Road) and calculate three different gates (Forget, Input, Output) at every single word, training an LSTM on a massive dataset can take weeks.

In 2014, researchers asked a simple question: "Do we really need all these moving parts?"

The result was the Gated Recurrent Unit (GRU).


The 2-in-1 Concept​

Think about buying shampoo and conditioner. An LSTM is like buying them in two separate bottles. It's highly customizable, but it takes more time in the shower and costs more money.

A GRU is like a 2-in-1 Shampoo and Conditioner. It merges things together to save time and computer memory!

1. Merging the Roads​

The GRU completely gets rid of the VIP Highway (Cell State). It merges the long-term memory and the short-term memory back into a single road (just the Hidden State).

2. Merging the Gates​

An LSTM has separate valves for "forgetting" old stuff and "inputting" new stuff. The GRU combines these into a single Update Gate.

The Update Gate: Instead of deciding what to forget and what to add separately, the Update Gate works like a sliding scale.

If it slides to the left (0), it completely ignores the new word and keeps 100% of its old memory. If it slides to the right (1), it completely throws away the old memory and replaces it with the new word!

It also uses a Reset Gate to decide how much of the past is relevant to calculating the current word.

Which one should you use?​

The million-dollar question: If GRUs are simpler and faster, are they better than LSTMs?

The answer is: It's a tie.

  • In 90% of real-world AI tasks, GRUs and LSTMs get the exact same accuracy.
  • GRUs train much faster because they have less math to do.
  • LSTMs are sometimes slightly better if you have a massive dataset with incredibly long and complex sentences, because the separate VIP Highway gives the AI a bit more fine-grained control.

Usually, developers will try a GRU first. If it works, great! If they need a little more power, they switch to an LSTM.

Next Up: We've solved the memory problem. But what if the AI needs to look into the future to understand the past? Let's bend time with Bidirectional Recurrence!